How a global platform team cut Kubernetes add-on upgrade cycles from weeks to days with agentic engineering
Twelve production EKS clusters. Forty-plus add-ons per upgrade. A team of eight replaced manual runbooks with governed AI agents — and turned a six-week slog into a 30-minute assessment.
From six weeks of archaeology to a single governed session
For a global platform team, managing the Kubernetes lifecycle was becoming an uphill battle. With dozens of add-ons per cluster and rapidly growing infrastructure, a single upgrade could consume up to six weeks of manual effort — days buried in release notes, fragile Helm forks, and validation checks that offered little certainty of success.
Seeking a way out, the team moved beyond simple AI assistance to a governed agentic engineering approach: fragmented, ad-hoc workflows replaced with a unified, intelligent lifecycle. By integrating live cluster inventory, versioned agent skills, and automated in-cluster validation, they built a system that doesn't just suggest changes — it orchestrates them with precision and governance.
The results were profound. What once took weeks of tedious research and manual execution is now a structured, multi-hour session — compressing upgrade cycles from weeks to days while restoring confidence with an auditable, repeatable playbook for enterprise-grade Kubernetes operations.
- 01Upgrade cycles dropped from 3–6 weeks of manual work to ~30 minutes of automated assessment plus a multi-hour agent session.
- 02The stack runs 26 allowlisted MCP tools, 4 versioned agent skills, and detection for 8 distinct IaC patterns.
- 03Every upgrade passes human-in-the-loop capability gates — cluster mutation and RBAC changes require explicit approval, off by default.
- 04Upgrades are validated in-cluster — resource health, API readiness, end-to-end checks — never assumed from pod status alone.
- 05Spans 12 production EKS clusters and 500+ engineers at a FinTech infrastructure company.
Success became an unsustainable operational burden
As the organization's infrastructure scaled, every minor EKS update triggered a complex chain of dependencies across more than 40 add-ons — with no centralized source of truth to manage it.
- 0140+ add-ons to assess for EOL status, compatibility, and breaking changes on every EKS minor bump
- 02Release notes scattered across GitHub, vendor docs, and internal wikis
- 03IaC spread across multiple patterns — forked Helm charts, Terraform helm_release, Argo CD Applications, raw manifests
- 04No single source of truth for what was actually installed vs. what Git declared
- 05Generic AI tools that could edit files but had no cluster context, no version control on instructions, and no validation framework
Attempts to use generic AI tools only highlighted the gap. While these assistants could edit code, they lacked the cluster context and validation frameworks needed for safe production changes — a complexity gap that led to delayed security adoptions and team burnout.
Agents engineered into the upgrade lifecycle, not bolted onto it
Rather than adding a chatbot to existing runbooks, the team built agents into every stage of the lifecycle — from observation to validation — across a multi-service platform.
Observe — live cluster inventory
A CronJob deployed into every cluster's platform namespace runs every 12 hours, collecting 60+ Kubernetes resource kinds plus all CRDs and custom resources, capturing cluster metadata (K8s version, cloud provider, node groups, environment labels), and uploading compressed manifest batches to a central ingestion API.
Know — always-on knowledge agents
A knowledge pipeline keeps upgrade intelligence current without manual curation:
Assess — upgrade lifecycle orchestration
When a platform engineer triggers an Upgrade Assessment for a target K8s version:
- A workflow engine starts a generation job (~30 min)
- The assessment engine reads cluster scan data, upgrade recommendations, and curated release notes
- It produces package install changes — one per add-on needing upgrade — with reasons and breaking-change previews
- Stages and steps surface in the upgrade console, including IDE deep-links that open the next agent session with context pre-loaded
Package — context bundles + MCP tooling
26 MCP tools and a context bundle format let AI agents work locally instead of stuffing the LLM window. The planning agent skill creates assessments, selects projects, generates context, and writes a handoff manifest for the next agent.
Execute — versioned agent skills
Semver agent skills — not one-off prompts — ship with MCP allowlists and capability gates:
The IaC upgrade skill detects 8 IaC patterns (forked charts, Terraform Helm, Argo CD, raw manifests, and more), builds a Values Delta Map, runs a three-way merge via the platform CLI, and enforces a 100% action-completeness gate before claiming success.
Validate — safety, health & readiness checks
Automated preflight/postflight validation is a first-class upgrade step: resource health (deployment/pod status), API readiness (webhook and API connectivity), and end-to-end checks that create test resources and verify real functionality — for example, confirming cert-manager can issue a certificate. The validation agent skill resolves the correct container image per project and version, surfaces RBAC for human approval, and runs checks as in-cluster Jobs.
The agentic upgrade lifecycle, visualized
Cycle time, compressed
What the platform runs on
Manual vs. agentic upgrade maturity
Qualitative index from platform architecture design, scored 1–5.
What "governed agentic" actually means in production
What's under the hood
Campaign proof points
Agent skill registry = package manager for AI behavior. Git → validate → publish → MCP bootstrap at runtime.
Proposer–Judge loops. Compatibility and upgrade-advisor agents separate AI proposals from AI verification before anything ships.
Context bundles beat context windows. Charts stay on disk; agents read plan metadata and merge locally.
100% completeness invariant. The IaC upgrade skill forbids partial success claims.
Six capability gates default off. Cluster mutation and RBAC changes require explicit human approval.
IDE deep-links. Upgrade steps open Cursor with assessment context pre-wired.
From observe to validate in one platform. No tool switching across the upgrade lifecycle.
- Upgrade Assessments & PlansOrchestration service
- In-Cluster Validation FrameworkValidation pack service
- Compatibility Knowledge GraphKnowledge agents
- Release Note CurationRelease notes pipeline
How long does a Kubernetes add-on upgrade normally take?
What is agentic engineering in a Kubernetes context?
What is an MCP tool?
Does agentic AI replace human approval in production upgrades?
What IaC patterns does the upgrade skill support?
Established in 2012, Xgrid has a history of delivering a wide range of intelligent and secure cloud infrastructure, user interface and user experience solutions. Our strength lies in our team and its ability to deliver end-to-end solutions using cutting edge technologies.
NAVIGATE
Cloud & DevOps Web & Mobile Apps Temporal Digital Marketing GTM Engineering Marketo Consulting HubSpot Consulting Company Careers ResourcesOFFICE ADDRESS
US Address:
Plug and Play Tech Center, 440 N Wolfe Rd, Sunnyvale, CA 94085
Dubai Address:
Dubai Silicon Oasis, DDP, Building A1, Dubai, United Arab Emirates
Pakistan Address:
Xgrid Solutions (Private) Limited, Bldg 96, GCC-11, Civic Center, Gulberg Greens, Islamabad
Xgrid Solutions (Pvt) Ltd, Daftarkhwan (One), Building #254/1, Sector G, Phase 5, DHA, Lahore